Semantic Web  •  Week 02

XML & XML Schema

Structured data representation, validation with XSD, and the XML model of the allergy data

Graduate Semantic Web Course  •  CMPE 583

Week 02  •  Objectives

By the end of this week

  • you will be able to tell well-formed XML from valid XML in practice;
  • you will be able to model the allergy data in XML and justify your element/attribute decisions;
  • you will be able to write datatype, facet, cardinality and key constraints in XSD;
  • you will be able to validate an XML document against a schema in Java;
  • you will be able to explain why the tree model of XML is not enough for the graph model of RDF.

Recap from last week

  • The Document Web encodes presentation; the Data Web encodes meaning.
  • The stack: IRI → XML → RDF → RDFS → OWL → rules/queries.
  • The risk chain in the project: Person → Product → FoodAdditives → Allergy.
  • This week we are on the second step of the stack: syntax and validation.

Output of this week

products.xml persons.xml allergy.xsd Validate.java

01

Solution of Assignment 1

Toolchain setup, ontology inventory and the class/individual distinction.

Solution 1.1 — Checking the environment

$ java -version openjdk version "1.8.0_392" $ mvn -v Apache Maven 3.9.6 Protégé 5.6.4 Tabs: Entities · Individuals SWRLTab · OntoGraf
CheckExpected evidence
Protégé startsVersion in the title bar
SWRLTab is visibleScreenshot of the tab
The OWL file opensThe class tree is populated
JDK / MavenVersion output

If SWRLTab is not visible: File → Check for plugins → install the SWRLTab and SWRLAPI plugins and restart Protégé.

Solution 1.2 — Inventory of ALLERGY_FIXED.owl

Entity typeCountExamples
Class6Person, Product, FoodAdditives, Allergy, Adult, PersonAtRisk
Object property9Contain, Triggers, hasAllergy, ChooseProduct, Effected_Allergen + 4 sub-properties
Data property6hasName, hasAge, hasWeight, hasHeight, hasBMI, hasProductName
Individual174 persons, 4 products, 5 additives, 4 allergies
SWRL rule7S1 … S7
Disjointness axiom1AllDisjointClasses (4 top-level classes)

The classes Adult and PersonAtRisk have no asserted members; membership comes from the rules.

Solution 1.3 — Class or individual?

TermCorrect modellingReason
Food additiveClass (FoodAdditives)A kind, it has members
NisinIndividualOne specific substance
Lactose allergyIndividual (Lactose)A specific member of the Allergy class
ProductClassIt covers barcoded individuals
Eti ChocolateIndividual + hasProductNameThe barcode is identity, the name is data
Person at riskClass, populated by a ruleMembership is computed, not asserted

Solution 1.4 — Modelling EAN_00005

EAN_00005 a Product ; Contain Whey_Protein , Wheat_Starch ; hasProductName "Protein Biscuit" . Whey_Protein a FoodAdditives ; Triggers Lactose . Wheat_Starch a FoodAdditives ; Triggers Gluten .

TC_004 (allergic to Egg and Gluten) chooses this product:

S6 → Effected_Allergen(TC_004, Wheat_Starch) S7 → PersonAtRisk(TC_004)

Whey_Protein triggers nothing here: TC_004 has no lactose allergy, so hasAllergy(?p, ?al) in the rule body does not bind.

02

XML Fundamentals

The tree model, well-formedness rules, namespaces.

What is XML?

  • A text format that marks data up with tags and is independent of any application.
  • There is no fixed tag set; the domain expert defines the tags.
  • It carries structure, not meaning — the meaning stays in the schema and in the application.
<?xml version="1.0" encoding="UTF-8"?> <product ean="EAN_00004"> <name>Eti Chocolate</name> <additives> <additive>Nisin</additive> </additives> </product>

Conditions for being well-formed

  1. There must be exactly one root element.
  2. Every opened tag must be closed.
  3. Nesting must be in the correct order.
  4. Tag names are case sensitive.
  5. Attribute values must be quoted.
<!-- HATALI --> <product ean=EAN_00004> ← no quotes <Name>Eti</name> ← case mismatch <additives><additive>Nisin </additives></additive> ← wrong nesting order </product>

No parser reads a document that is not well-formed — it is the minimum condition before validation.

The parts of an XML document

<?xml version="1.0" encoding="UTF-8"?> ← prolog <!-- Product catalogue of the allergy project --> ← comment <products count="4"> ← root element + attribute <product ean="EAN_00003"> ← child element <name>Dardanel Ton</name> ← text (PCDATA) <additive ref="Casein"/> ← empty element </product> </products>

A document is a tree: every node has exactly one parent. This restriction is why we move to RDF in Week 03.

Element or attribute?

Attribute-heavy

<person tc="TC_001" name="Ayse" age="38" weight="67.5" height="1.68"/>

Short; but it cannot repeat and cannot carry structure.

Element-heavy

<person tc="TC_001"> <name>Ayse</name> <age>38</age> <allergy ref="Lactose"/> <allergy ref="Fish"/> </person>

multiple values and extension are not possible.

Rule of thumb: identity and metadata as attributes, field data as elements. A person may have more than one allergy, so allergy must be an element.

Namespaces: same name, different meaning

<cat:products xmlns:cat="http://EMU/catalog#" xmlns:med="http://EMU/medical#"> <cat:product ean="EAN_00003"> <cat:name>Dardanel Ton</cat:name> <med:risk level="high"/> </cat:product> </cat:products>
  • xmlns:prefix="IRI" declarations resolve name clashes.
  • It is the IRI that carries the meaning, not the prefix.
  • A declaration without a prefix (xmlns=) sets the default namespace.
  • The OWL file of the project uses the same mechanism: rdf:, owl:, swrl:.

The XML header of the project

<rdf:RDF xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:xsd="http://www.w3.org/2001/XMLSchema#" xmlns:rdfs="http://www.w3.org/2000/01/rdf-schema#" xmlns:owl="http://www.w3.org/2002/07/owl#" xml:base="http://EMU/AllergyOntology" xmlns="http://EMU/AllergyOntology#" xmlns:swrl="http://www.w3.org/2003/11/swrl#">
PrefixRole
rdf, rdfsGraph and vocabulary syntax
owlOntology constructs
xsdDatatypes (int, double, string)
swrlRule axioms
(default)The project's own terms: Person, Nisin …

Special characters and CDATA

CharacterEntity
<&lt;
>&gt;
&&amp;
"&quot;
'&apos;
<ingredients> Contains milk &amp; cocoa </ingredients> <label><![CDATA[ E322 (soya lesitini) < 0.5% ]]></label>

Percent signs and "&" are common on food labels; this is the most frequent source of parsing errors.

Encoding: pitfalls with non-ASCII characters

<?xml version="1.0" encoding="UTF-8"?> <product ean="EAN_00006"> <name>Ulker Chocolate Wafer</name> <note>Contains milk · 30% cocoa</note> </product> <!-- if saved as ISO-8859-9: --> <!-- Ülker Çikolatalı -->
PitfallResult
Declared UTF-8, file saved as ANSIParsing error
UTF-8 with BOM"Content is not allowed in prolog"
Dotless-i case mappingUse Locale.ROOT in Java
Non-ASCII letter in an IRIKeep local names in ASCII

Project rule: label text may be in any language, but the local name of an IRI is always ASCII (Soy_Lecitin, label: "Soy lecithin").

The tree model of the document

products ├── product @ean=EAN_00003 │ ├── name "Dardanel Ton" │ ├── additive @ref=Casein │ └── additive @ref=Sodium_Ascorbite └── product @ean=EAN_00004 ├── name "Eti Chocolate" ├── additive @ref=Nisin └── additive @ref=Soy_Lecitin

This tree does not say which allergy an additive triggers; we point to another document with ref — the link is made by the application.

Example: products.xml

<?xml version="1.0" encoding="UTF-8"?> <products xmlns="http://EMU/allergy/catalog" count="4"> <product ean="EAN_00001"> <name>ETI Cracker</name> <additive ref="Alginic_Acid"/> </product> <product ean="EAN_00002"> <name>Ulker Damak</name> <additive ref="Phospore"/> <additive ref="Soy_Lecitin"/> </product>
<product ean="EAN_00003"> <name>Dardanel Ton</name> <additive ref="Casein"/> <additive ref="Sodium_Ascorbite"/> </product> <product ean="EAN_00004"> <name>Eti Chocolate</name> <additive ref="Ascorbic_Acid"/> <additive ref="Nisin"/> <additive ref="Soy_Lecitin"/> </product> </products>

Additives point to a separate document with ref ; there is no direct link between a product and an allergy.

Example: additives.xml

<additives xmlns="http://EMU/allergy/catalog"> <additive id="Nisin" code="E234"> <label>Nisin</label> <triggers allergy="Lactose"/> </additive> <additive id="Casein" code="E290"> <label>Casein</label> <triggers allergy="Lactose"/> </additive> <additive id="Soy_Lecitin" code="E322"> <label>Soy Lecithin</label> <triggers allergy="Egg"/> </additive> <additive id="Ascorbic_Acid" code="E300"> <label>Ascorbic Acid</label> </additive> </additives>

Ascorbic_Acid has no triggers — exactly the situation in the ontology.

In XML this "gap" is just a missing value. In OWL, under the Open World Assumption, it means "unknown" — the same data, two different readings.

Example: persons.xml

<persons xmlns="http://EMU/allergy/profile"> <person tc="TC_001"> <name>Ayse</name> <age>38</age> <weight unit="kg">67.5</weight> <height unit="m">1.68</height> <allergy ref="Lactose"/> <choice ean="EAN_00004"/> </person>
<person tc="TC_003"> <name>MEHMET</name> <age>35</age> <weight unit="kg">93.0</weight> <height unit="m">1.87</height> <allergy ref="Fish"/> <allergy ref="Lactose"/> <choice ean="EAN_00003"/> </person> </persons>

BMI is absent here: computed values are not kept in the source data — rule S4 will produce it in Week 07.

03

Validation: DTD and XML Schema

Datatypes, constraints, keys and reading error messages.

The difference between well-formed and valid

AspectWell-formedValid
What it checksSyntaxConformance to the schema
ReferenceThe XML 1.0 rulesThe DTD / XSD document
"age=abc"ValidError: int expected
Unknown elementValidError
Required field missingValidError

Label data comes from an external source, so validation is compulsory: dirty data in the ontology produces wrong risk inferences.

Validation with a DTD, and its limits

<!ELEMENT products (product+)> <!ELEMENT product (name, additive*)> <!ATTLIST product ean ID #REQUIRED> <!ELEMENT name (#PCDATA)> <!ELEMENT additive EMPTY> <!ATTLIST additive ref IDREF #REQUIRED>
  • No datatypes: age may be "abc".
  • No namespace support.
  • Numeric ranges, patterns and decimal constraints cannot be expressed.
  • It is not XML itself, so tooling is harder.

This is why we use XSD in the project; we learn DTD only to read legacy documents.

The skeleton of an XML Schema document

<?xml version="1.0" encoding="UTF-8"?> <xs:schema xmlns:xs="http://www.w3.org/2001/XMLSchema" xmlns="http://EMU/allergy/catalog" targetNamespace="http://EMU/allergy/catalog" elementFormDefault="qualified"> <xs:element name="products" type="ProductsType"/> <!-- type definitions go here --> </xs:schema>

targetNamespace states which namespace the schema defines; the document must use the same namespace.

Built-in datatypes

XSD typeExample valueField in the project
xs:string"Eti Chocolate"hasName, hasProductName
xs:int38hasAge
xs:double67.5hasWeight, hasHeight, hasBMI
xs:booleantruelabel validated?
xs:date2026-09-08expiry date
xs:ID / xs:IDREFEAN_00004barcode and reference

The same type names appear on the OWL side: rdfs:range = &xsd;double. XSD datatypes are the shared vocabulary of the two worlds.

Define your own type: facets

<xs:simpleType name="AgeType"> <xs:restriction base="xs:int"> <xs:minInclusive value="0"/> <xs:maxInclusive value="120"/> </xs:restriction> </xs:simpleType> <xs:simpleType name="WeightType"> <xs:restriction base="xs:double"> <xs:minExclusive value="0"/> <xs:maxInclusive value="400"/> </xs:restriction> </xs:simpleType>
FacetWhat it does
minInclusiveLower bound (inclusive)
maxExclusiveUpper bound (exclusive)
lengthLength
patternRegular expression
enumerationSet of allowed values
fractionDigitsDecimal digits

Enumeration: allergy types

<xs:simpleType name="AllergyType"> <xs:restriction base="xs:string"> <xs:enumeration value="Lactose"/> <xs:enumeration value="Gluten"/> <xs:enumeration value="Egg"/> <xs:enumeration value="Fish"/> </xs:restriction> </xs:simpleType>

Here XSD builds a closed world: a value outside the list is an error.

The OWL counterpart owl:oneOf can express this; in the project, however, we keep allergies as individuals, because adding a new allergy type should not force a schema change.

Pattern: barcode and identifier format

<xs:simpleType name="EanType"> <xs:restriction base="xs:string"> <xs:pattern value="EAN_[0-9]{5}"/> </xs:restriction> </xs:simpleType> <xs:simpleType name="TcType"> <xs:restriction base="xs:string"> <xs:pattern value="TC_[0-9]{3}"/> </xs:restriction> </xs:simpleType>
ValueEanType result
EAN_00004Valid
EAN_4Error — five digits required
ean_00004Error — upper case required

Complex type: product

<xs:complexType name="ProductType"> <xs:sequence> <xs:element name="name" type="xs:string"/> <xs:element name="additive" type="AdditiveRefType" minOccurs="0" maxOccurs="unbounded"/> </xs:sequence> <xs:attribute name="ean" type="EanType" use="required"/> </xs:complexType> <xs:complexType name="AdditiveRefType"> <xs:attribute name="ref" type="xs:string" use="required"/> </xs:complexType>

minOccurs="0": a product with no declared additive is still valid — a missing value does not break the schema.

sequence, choice, all

ConstructMeaningAllergy example
sequenceAll of them, in the given ordername, then the additive list
choiceOnly one of themya ean ya internalCode as identifier
allAll of them, order freeweight, height, age
groupA reusable groupthe measurement block
<xs:choice> <xs:element name="ean" type="EanType"/> <xs:element name="internalCode" type="xs:string"/> </xs:choice>

Cardinality constraints

DeclarationMeaning
minOccurs="1"Required (default)
minOccurs="0"Optional
maxOccurs="unbounded"Unbounded repetition
use="required"Required attribute

The OWL counterpart

Product ⊑ ≥1 Contain.FoodAdditives

An XSD constraint rejects data; an OWL restriction produces an inference. The same sentence, two different behaviours — covered in detail in Week 05.

key and keyref: referential integrity

<xs:element name="catalog" type="CatalogType"> <xs:key name="additiveKey"> <xs:selector xpath="additives/additive"/> <xs:field xpath="@id"/> </xs:key> <xs:keyref name="additiveRef" refer="additiveKey"> <xs:selector xpath="products/product/additive"/> <xs:field xpath="@ref"/> </xs:keyref> </xs:element>

This way a typo such as <additive ref="Nisiin"/> is caught during validation — a wrong individual never reaches the ontology.

The full schema: allergy.xsd

<xs:schema xmlns:xs="http://www.w3.org/2001/XMLSchema" targetNamespace="http://EMU/allergy/catalog" xmlns="http://EMU/allergy/catalog" elementFormDefault="qualified"> <xs:element name="products"> <xs:complexType> <xs:sequence> <xs:element name="product" type="ProductType" maxOccurs="unbounded"/> </xs:sequence> <xs:attribute name="count" type="xs:int"/> </xs:complexType> </xs:element>
<xs:complexType name="ProductType"> <xs:sequence> <xs:element name="name" type="xs:string"/> <xs:element name="additive" type="AdditiveRefType" minOccurs="0" maxOccurs="unbounded"/> </xs:sequence> <xs:attribute name="ean" type="EanType" use="required"/> </xs:complexType> <xs:simpleType name="EanType"> <xs:restriction base="xs:string"> <xs:pattern value="EAN_[0-9]{5}"/> </xs:restriction> </xs:simpleType> </xs:schema>

Reading validation errors

<product ean="EAN_4">
cvc-pattern-valid: 'EAN_4' is not facet-valid with respect to pattern 'EAN_[0-9]{5}'
<age>abc</age>
cvc-datatype-valid.1.2.1: 'abc' is not a valid value for 'int'
<product> <!-- no ean -->
cvc-complex-type.4: Attribute 'ean' must appear on element 'product'

The cvc-* code at the start of the message tells you which rule was violated; it is the fastest way to debug.

Validation in Java: Validate.java

import javax.xml.XMLConstants; import javax.xml.validation.*; import org.xml.sax.SAXException; import java.io.File; public class Validate { public static void main(String[] a) throws Exception { SchemaFactory sf = SchemaFactory.newInstance( XMLConstants.W3C_XML_SCHEMA_NS_URI); Schema schema = sf.newSchema(new File("allergy.xsd")); Validator v = schema.newValidator(); try { v.validate(new javax.xml.transform.stream.StreamSource( new File("products.xml"))); System.out.println("products.xml GECERLI"); } catch (SAXException e) { System.out.println("HATA: " + e.getMessage()); } } }

Choosing a parser: DOM, SAX, StAX

ModelApproachMemoryIn the allergy project
DOMLoads the document as a treeHighSmall catalogue, navigation with XPath
SAXEvent based, single passVery lowA store dump with thousands of products
StAXPull-based streamingLowPartial reading, early exit
JAXBObject mappingMediumBinding to the Product/Person classes

Rule of thumb: DOM if the data fits in memory and you need random access; SAX or StAX if large data arrives as a stream.

Java + DOM: reading the catalogue

DocumentBuilderFactory f = DocumentBuilderFactory.newInstance(); f.setNamespaceAware(true); Document doc = f.newDocumentBuilder() .parse(new File("products.xml")); NodeList ps = doc.getElementsByTagNameNS( "http://EMU/allergy/catalog", "product"); for (int i = 0; i < ps.getLength(); i++) { Element p = (Element) ps.item(i); String ean = p.getAttribute("ean"); NodeList as = p.getElementsByTagNameNS("*","additive"); for (int j = 0; j < as.getLength(); j++) System.out.println(ean + " Contain " + ((Element) as.item(j)).getAttribute("ref")); }
EAN_00001 Contain Alginic_Acid EAN_00002 Contain Phospore EAN_00002 Contain Soy_Lecitin EAN_00003 Contain Casein EAN_00003 Contain Sodium_Ascorbite EAN_00004 Contain Ascorbic_Acid EAN_00004 Contain Nisin EAN_00004 Contain Soy_Lecitin

This output is already in triple form — in Week 08 the same lines become axioms through the OWL API.

Java + SAX: reading as a stream

SAXParserFactory.newInstance().newSAXParser().parse( new File("market_dump.xml"), new DefaultHandler() { String ean; public void startElement(String u, String l, String q, Attributes at) { if ("product".equals(l)) ean = at.getValue("ean"); if ("additive".equals(l)) risk(ean, at.getValue("ref")); } }); // risk(): if the additive triggers Lactose, add it to the warning list

SAX keeps nothing in memory: a 500 MB store dump can be scanned with constant memory; in exchange you cannot go back and navigate.

Selecting data with XPath

XPathResult
/products/product/@eanAll barcodes
//product[additive/@ref='Nisin']/nameNames of products containing Nisin
count(//product[@ean='EAN_00004']/additive)3
//person[age>=18]/nameNames of adults
//person[allergy/@ref='Lactose']/@tcTC_001, TC_002, TC_003

Note: the last query finds only asserted allergies. The question "who is at risk" cannot be answered in XPath, because the additive → allergy chain lies outside the document.

A report with XQuery: which product affects whom?

for $p in doc("persons.xml")//person let $ean := $p/choice/@ean let $prod := doc("products.xml") //product[@ean = $ean] for $a in $prod/additive/@ref let $trg := doc("additives.xml") //additive[@id = $a]/triggers/@allergy where $trg = $p/allergy/@ref return <risk tc="{$p/@tc}" ean="{$ean}" additive="{$a}"/>
<risk tc="TC_001" ean="EAN_00004" additive="Nisin"/> <risk tc="TC_002" ean="EAN_00003" additive="Casein"/> <risk tc="TC_003" ean="EAN_00003" additive="Casein"/> <risk tc="TC_003" ean="EAN_00003" additive="Sodium_Ascorbite"/>

We will write the same result in a single SWRL rule (S6). The difference: here you build the chain; there the engine applies the rule and writes the result permanently into the ontology.

From XML to RDF/XML with XSLT

<xsl:template match="product"> <owl:NamedIndividual rdf:about="#{@ean}"> <rdf:type rdf:resource="#Product"/> <hasProductName> <xsl:value-of select="name"/> </hasProductName> <xsl:for-each select="additive"> <Contain rdf:resource="#{@ref}"/> </xsl:for-each> </owl:NamedIndividual> </xsl:template>
<!-- output --> <owl:NamedIndividual rdf:about="#EAN_00004"> <rdf:type rdf:resource="#Product"/> <hasProductName>Eti Chocolate </hasProductName> <Contain rdf:resource="#Ascorbic_Acid"/> <Contain rdf:resource="#Nisin"/> <Contain rdf:resource="#Soy_Lecitin"/> </owl:NamedIndividual>

Using this transformation instead of typing label data into the ontology by hand is the most practical way to feed the project with real data.

04

Is XML Enough?

The difference between a tree and a graph, and why we move to RDF.

The tree model and the graph model

XML: a tree

person └── allergy @ref="Lactose" (the link is only a name)

The reference is text; only the application code knows its meaning.

RDF: a graph

TC_001 hasAllergy Lactose . Nisin Triggers Lactose . EAN_00004 Contain Nisin .

Nodes are shared; inference runs along the chain.

Why is XML alone not enough?

What is missingConsequenceLayer that solves it
Order carries meaningSame information, different treeRDF (unordered triples)
Links are unnamedMeaning lives in the codeRDF predicates
No class / subclassHierarchy in the codeRDFS
No constraints or logicContradictions cannot be foundOWL
No inferenceOnly what is written is knownReasoner + SWRL
No global identityData merging by handIRI

XML remains indispensable: our RDF/XML and OWL files are XML — it stays as the transport layer.

The same product, two representations

products.xml

<product ean="EAN_00003"> <name>Dardanel Ton</name> <additive ref="Casein"/> <additive ref="Sodium_Ascorbite"/> </product>

ALLERGY_FIXED.owl

<owl:NamedIndividual rdf:about="#EAN_00003"> <rdf:type rdf:resource="#Product"/> <Contain rdf:resource="#Casein"/> <Contain rdf:resource="#Sodium_Ascorbite"/> <hasProductName rdf:datatype="&xsd;string"> Dardanel Ton</hasProductName> </owl:NamedIndividual>

The syntax is almost identical; the difference is the globally identified link built with rdf:resource. That difference is exactly the topic of Week 03.

Checklist for catalogue data

  • Every document declares a namespace; use the prefix consistently.
  • Identifiers are ASCII and pattern constrained (EAN_[0-9]{5}).
  • Carry the unit of measure as an attribute (unit="kg").
  • Do not keep a computed value (BMI) in the source data.
  • Keep the additive → allergy mapping in exactly one place.
  • Validate with XSD before every load; log the error.
  • Save as UTF-8 without a BOM.
  • Keep the transformation (XSLT) under version control; never fix output by hand.

This list is applied directly in Week 08, when the ontology is fed with real catalogue data.

05

Assignment and Project Step

Build your own product catalogue with XML + XSD.

Assignment 2 — Catalogue XML and its schema

  1. Write products.xml for five packaged products of your own choice.
  2. Define the additive → allergy mapping in additives.xml.
  3. Write allergy.xsd with a barcode pattern, an age range and an allergy enumeration.
  4. Enforce referential integrity with keyref.
  5. Produce two invalid documents and report the validation messages.

Deliverable

Four files + a 2-page report: design decisions (element vs attribute), the reasons for your facets, and an interpretation of the error messages.

We will discuss the solution at the start of Week 03.

Assessment criteria

CriterionWeightExpected
Well-formed + valid documents25%Validation passes without errors
Expressiveness of the schema30%Facets, patterns, enumeration and cardinality are used
Referential integrity20%key/keyref works
Design rationale15%Element/attribute decisions are defended
Error analysis10%cvc codes are interpreted correctly

References

  • W3C — Extensible Markup Language (XML) 1.0, 5th Edition.
  • W3C — XML Schema Part 0: Primer; Part 1: Structures; Part 2: Datatypes.
  • W3C — Namespaces in XML 1.0; XML Path Language (XPath) 3.1.
  • Harold, E. R., Means, W. S. — XML in a Nutshell, O'Reilly.
  • Allemang & Hendler — Semantic Web for the Working Ontologist, Part 3 (moving to RDF).

Summary  •  1 / 2

XML and validation

  • XML carries structure, not meaning; the meaning stays in the schema and the application.
  • Well-formed is the minimum condition; valid means conforming to the schema.
  • XSD gives datatypes, facets, cardinality and key constraints; DTD does not.
  • Identity and metadata become attributes; repeatable field data becomes elements.
  • Validation is the first line of defence against dirty data entering the ontology.

Summary  •  2 / 2

Contribution to the project and the next step

  • The allergy data was split into three documents: products, additives, persons.
  • Barcode and identifier formats were secured with patterns.
  • The skeleton of the XML → RDF/XML transformation was built with XSLT.
  • XPath queries asserted data; it cannot query the risk chain.

In Week 03

RDF & RDFS: the triple model, thinking in graphs, Turtle syntax and building the allergy vocabulary.

Also: the detailed solution of Assignment 2.

Review Questions

Test yourself

  1. Can a document be well-formed but not valid? Give an example.
  2. Why did we model the allergy list as elements rather than attributes?
  3. minOccurs="0" and the ≥1 restriction in OWL — what is the behavioural difference?
  4. Which error does keyref catch, and which one does it not?
  1. Why can the "persons at risk" query not be written in XPath?
  2. Explain the relation between a namespace prefix and an IRI.
  3. What is the difference in world assumption between XSD enumeration and OWL oneOf?
  4. In the XSLT transformation, why was rdf:resource used instead of a text value?

Exercise  •  In class

Test the schema against an invalid document

The document below contains three validation errors. Find them, predict the cvc message and fix them.

<products count="two"> <product ean="EAN_105"> <additive ref="Nisin"/> <name>Test Bar</name> </product> </products>

Hint

Think in order: attribute type, pattern conformance, element order inside xs:sequence.

The solution comes in the assignment-solution part of Week 03.